Back

Nature Methods

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Nature Methods's content profile, based on 385 papers previously published here. The average preprint has a 0.38% match score for this journal, so anything above that is already an above-average fit.

1
Fault-tolerant 3D reconstruction from 2D spatial proteomics sections

Zhang, Z.; Tan, Y.; Nolan, G.; Snyder, M.; Ma, Z.

2026-06-28 bioinformatics 10.64898/2026.06.23.733649 medRxiv
Top 0.1%
45.4%
Show abstract

Reconstructing 3D molecular volumes from sparsely sampled 2D tissue sections is limited by per-section marker dropout and tissue loss. We present 3D-Omics-Flow, a generative pipeline that jointly repairs damaged sections and interpolates between them at single-cell resolution. Across datasets spanning health and disease, 3D-Omics-Flow expands 3D spatial proteomics to practical sampling regimes, enabling atlas construction and downstream analysis from imperfect 2D section stacks.

2
A refined Saccharomyces cerevisiae reference transcriptome from Direct RNA Sequencing, with a reusable pipeline for UTR annotation updates

Rossini, O.; Cleynen, A.; Shirokikh, N. E.

2026-06-30 genomics 10.64898/2026.06.30.735557 medRxiv
Top 0.1%
42.1%
Show abstract

Untranslated regions (UTRs) flanking the coding sequence govern mRNA translation, localisation, stability, and decay, making accurate UTR boundaries essential for quantitative RNA sequencing and the study of post-transcriptional control in Saccharomyces cerevisiae and beyond. Reference transcriptomes built from short-read sequencing have been invaluable to the yeast community, yet in a genome as gene-dense as that of S. cerevisiae, short reads frequently cannot be assigned to a single transcript of origin, leaving roughly one quarter of transcripts without a confidently defined UTR. Here we use Oxford Nanopore Direct RNA Sequencing (DRS), in which each full-length polyadenylated molecule is read end to end, to resolve this ambiguity and deliver two complementary resources. First, an updated, ready-to-use S. cerevisiae S288C reference: change-point segmentation of per-gene DRS coverage defined boundaries for 5,416 of the 6,695 annotated genes, and a merge-max rule retaining the longer UTR from each source ensures no gene loses existing annotation. The result adds previously absent UTRs to 927 (5') and 896 (3') genes and extends 29.4% of 5' and 26.1% of 3' boundaries among comparable genes. Second, the complete, documented pipeline so that any laboratory can rebuild or update a transcriptome from its own DRS data. Validation on two independent datasets shows improved mapping rates, reduced soft-clipping, and metagene profiles consistent with genuine transcript signal.

3
RAEM: random-access electron microscopy for revisitable 3D imaging

Chandok, I. S.; Patel, M.; Wu, Y.; Berger, D.; Schalek, R.; Lichtman, J. W.; Samuel, A. D.; Meirovitch, Y.

2026-06-23 neuroscience 10.64898/2026.06.18.732873 medRxiv
Top 0.1%
39.9%
Show abstract

Volume electron microscopy is essential for understanding cells, tissues, and neural circuits in their native 3D context, but many biological specimens are too large to image exhaustively at nanometer resolution. Researchers therefore must choose between broad anatomical context and ultrastructural detail. We introduce random-access electron microscopy (RAEM), a framework for studying fixed tissue repeatedly across scales rather than imaging it once at a single resolution. RAEM first builds a lower resolution 3D survey of the specimen, then uses accumulated human or AI-derived knowledge of that volume to guide the microscope back to selected physical sites for high resolution imaging. By linking reconstructed 3D coordinates to precise electron-beam positions on the original sections, RAEM enables targeted imaging of membranes, vesicles, and other nanoscale structures within specimens that would be impractical to image exhaustively. We demonstrate RAEM with vesicle-resolved imaging of synaptic boutons in human cortex, targeted imaging of more than one million human cortical mitochondria, hierarchical imaging of a nematode nervous system, and retrospective targeting of a previously published petabyte-scale human cortical volume. RAEM turns serial-section EM into a query-driven, multi-resolution approach for scalable biomedical discovery.

4
Penumbria: Advanced 3D cell segmentation for biomedical imaging

Stockert, L.; Donovan, J.; Baier, H.

2026-07-01 bioinformatics 10.64898/2026.06.30.735527 medRxiv
Top 0.1%
39.5%
Show abstract

Quantitative analysis of three-dimensional cellular architecture is fundamental to understanding tissue organization, disease progression, and drug response. Yet 3D cell segmentation remains a critical bottleneck due to diverse cell morphologies, low signal-to-noise ratios, and data scarcity. We introduce Penumbria, a general-purpose 3D cell segmentation framework that achieves state-of-the-art accuracy across morphologically distinct cell populations and imaging conditions in volumetric microscopy. Penumbria formulates segmentation as a regression problem on distances to cell boundaries, supporting instance reconstruction without shape priors and permitting end-to-end GPU inference. A U-Net-based architecture with xLSTM bottleneck blocks and patch embeddings enables multi-scale feature extraction, long-range modeling of spatial context, and convolutional feature-volume tokenization. The model is extended with two modules: a Global Zernike Phase Layer, which learns Zernike-parameterized phase corrections in the frequency domain to undo optical aberrations such as defocus and tilt, and a Scaled Geocaps Layer, which samples features at fixed grid locations across multiple spatial scales, routing evidence between them such that a detection is only confident where concordance holds across scales simultaneously. Across four diverse 3D datasets selected to probe the limits of existing methods, Penumbria outperforms Cellpose-SAM across all evaluation thresholds and surpasses StarDist-3D on most datasets while matching it on Parhyale hawaiensis. Trained entirely from scratch, Penumbria achieves up to a 38% improvement in mean average precision over the second-best method. Strong boundary accuracy further supports downstream analyses such as quantifying membrane dynamics or protein localization.

5
A deep-learning-informed prior and Bayesian model for differential AP-MS interactome analysis

Seefelder, M.

2026-07-09 bioinformatics 10.64898/2026.07.06.736690 medRxiv
Top 0.1%
35.8%
Show abstract

Affinity-purification mass spectrometry (AP-MS) maps a bait proteins partners, but every purification also captures abundant non-specific background that masks genuine interactors. Established tools such as SAINTexpress and CompPASS score one evidence type, treat correlated signals as independent, and ignore prior knowledge of likely interactions. BayesInteractomics, an open-source Julia framework, addresses both limitations by combining machine learning with Bayesian statistics. A neural network trained on protein structures predicts direct binding. A calibrated meta-learner turns this into an informed prior. The prior guides a Bayesian copula-mixture model integrating three AP-MS evidence streams: enrichment, co-abundance, and detection reproducibility. Each candidate receives an interaction probability at a controlled false-discovery rate, optionally updated by structural docking. On synthetic data it ranks first in every benchmark (median AUROC 0.747), and across independent studies it raises high-confidence-call reproducibility from 21% to 79%. It also identifies which interactions are gained or lost between two conditions, unlike established tools.

6
HERO: A hierarchy-aware analysis pipeline for reducing and refining whole-brain atlas-mapped cellular datasets

Shipman, A. L.; Centanni, S. W.

2026-07-08 neuroscience 10.64898/2026.07.02.736093 medRxiv
Top 0.1%
31.3%
Show abstract

Advances in high-throughput mesoscale microscopy and machine learning-based image analysis pipelines have made unbiased whole-brain imaging widely accessible. However, translating the resulting atlas-mapped datasets into biologically meaningful results remains a substantial barrier owing to their sheer magnitude and complex hierarchical organization. Consequently, reporting structure and analysis methods vary widely across studies, under-mining rigor and reproducibility. To address this, we developed a user-friendly data reduction workflow, HERO (Hierarchy-aware Expression Region Organization), designed to perform hierarchy-aware selection, refinement, ranking, and visualization of whole-brain cell detec-tion datasets. The workflow is customizable to specific needs, requires minimal coding expe-rience, and outputs transparent, curated results. HERO is designed to function as a seamless plug-in within larger-scale whole-brain cell-detection analysis pipelines, providing efficient, unbiased region selection to streamline subsequent statistical analyses and comparative evaluations. Although HERO is developed with mouse cell-detection datasets, it can, in prin-ciple, be applied to any atlas-mapped dataset that contains hierarchical information. In sum, HERO offers a standardized analysis workflow to reduce whole-brain cell-detection datasets, transforming raw regional cell counts into curated results and advancing the effectiveness, interpretability, and accessibility of whole-brain imaging in neuroscience.

7
SSUplex: fast, both-strand extraction and origin-sorting of small-subunit rRNA for environmental DNA metabarcoding

O'Brien, A.; Vargas, J.; Acuna, I.; Parada, P.

2026-07-05 bioinformatics 10.64898/2026.07.02.736232 medRxiv
Top 0.1%
30.6%
Show abstract

Ribosomal RNA metabarcoding sits at the centre of how we characterise microbial and eukaryotic communities in environmental samples, and long-read sequencing has made full-length small-subunit (SSU; 16S/18S) profiling routine. The broadly conserved primers that make rRNA such a convenient marker are also its liability: by design they co-amplify organellar (mitochondrial, chloroplast) and cross-domain SSU alongside the intended target. Left unsorted before taxonomic assignment, these passengers are systematically misclassified, and the error propagates straight into estimates of community composition and diversity. Reads must therefore be detected, extracted, and sorted by origin before they ever reach a classifier. We present SSUplex, an open-source tool that detects SSU rRNA, assigns each read to one of five origins (bacteria, archaea, eukaryota, mitochondria, chloroplast), and extracts the SSU region for downstream classification. SSUplex reimplements the extraction-and-origin logic of the widely used Metaxa2 in the Rust programming language, scans both strands, and ships as a single dependency-light binary suited to long-read (Oxford Nanopore, PacBio HiFi) and short-read data. Benchmarked against Metaxa2 on public data, SSUplex reproduces Metaxa2 origin calls on full-length reads (96.8% concordance) and matches its extraction speed on small inputs, then pulls away to run up to approximately 3.4-fold faster with approximately 35% lower peak memory at 200,000 reads, the per-sample scale a long-read amplicon run typically reaches. We characterise a genuine, measured trade-off in the origin-ranking statistic, and we identify the bacteria-versus-mitochondria boundary as the method's one intrinsically lower-confidence edge. For the now-common workflow in which origin-sorted reads are handed to a dedicated classifier rather than classified in place, SSUplex is a fast, reproducible, embeddable stand-in for Metaxa2's extraction role. Source code and a benchmark harness that regenerates every result from public data are available under the MIT license at https://github.com/ayobi/ssuplex.

8
Spectral Unmixing: A modular and reproducible Python package for directed and blind spectral unmixing in multidimensional microscopy stacks

Musacchio, F.; Fuhrmann, M.

2026-07-10 neuroscience 10.64898/2026.07.06.736825 medRxiv
Top 0.2%
27.2%
Show abstract

Spectral bleed-through remains a persistent practical problem in multichannel fluorescence microscopy. Signal from one fluorophore can be recorded in the detection channel of another, thereby biasing intensity measurements, inflating apparent colocalization, and complicating the interpretation of dynamic microscopy data. Although many correction strategies exist, routine workflows often remain fragmented across ad hoc scripts, manually tuned graphical procedures, or method-specific blind-unmixing implementations with limited provenance. Here we present spectral-unmixing, an open-source Python package for reproducible linear spectral unmixing in multidimensional microscopy stacks. The package unifies directed two-channel correction with multiple alpha-estimation strategies, optional bidirectional two-channel correction through explicit inversion of a 2 x 2 mixing model, and PICASSO-family blind unmixing for multichannel data. Microscopy inputs are normalized at the API boundary to canonical TZCY X stacks, allowing the same unmixing code to be applied across file formats without manual axis handling. Machine-readable sidecar reports preserve the effective processing configuration and estimated coefficients for every output, so that workflows can be audited and reproduced. Synthetic and real-data-derived benchmarks show that the implemented workflows accurately estimate and correct bleed-through when their model assumptions are satisfied. In fixed-alpha two-channel simulations, the mean-ratio and linear-fit estimators recovered {approx} 0.283 for a ground-truth value of 0.28 and reduced target-channel normalized root mean squared error from approximately 0.029 to 0.003. In time-varying simulations, per-time-point estimation tracked coefficient drift substantially better than reference-time-point estimation. Bidirectional inversion recovered reciprocally mixed channels accurately when coefficients were known or well estimated. PICASSO-family benchmarks further showed a practical trade-off between reducing residual inter-channel dependence and preserving fluorophore identity, with MATLAB-style workflows behaving more conservatively and source-sink formulations providing stronger dependence suppression when meaningful directional priors are available. Together, these elements make spectral-unmixing a practical, transparent, and extensible platform for reproducible spectral unmixing of fluorescence microscopy data in neuroscience and other quantitative bioimage-analysis settings.

9
A High-Confidence Atlas of Protein Methylation Enables AI-Driven Detection of Methylated Peptides

Wang, S.; Hartmaring, Y.; Schlaffner, C. N.; Bowler-Barnett, E.; Martin, M.; Fan, J.; Sun, Z.; Renard, B. Y.; Jones, A. R.; Vizcaino, J. A. R.

2026-07-04 bioinformatics 10.64898/2026.07.01.733993 medRxiv
Top 0.2%
27.1%
Show abstract

Lysine and arginine methylation regulate chromatin dynamics, transcription, and cellular signaling, however confident mass spectrometry (MS)-based detection and localization of this modification remain challenging. We reanalyzed eight public human methylation-enriched datasets using an open and standardized workflow that integrates database searching via the Trans-Proteomic Pipeline with a decoy-based statistical method for the independent estimation of false localization rates (FLR). This yielded a high-confidence Human Methylation Atlas of 1,828 sites (57 methyl-lysine, 1,771 methyl-arginine) across 1,021 proteins, classified into Gold, Silver, and Bronze confidence tiers. This is far fewer sites than reported in previous studies, reflecting the application of stringent FLR control, and what we hypothesise is potential high-false discovery in previous analyses. We then leveraged this resource to adapt a deep learning-based methodology for the improved detection of methylated peptides. Three mouse methylation-enriched datasets were reanalysed to augment training and the phosphoproteomics-trained AHLF (ad hoc learning of peptide fragmentation) model was fine-tuned by transfer learning to create AHLF-Methylation. The model achieved mean ROC-AUC values of 0.824 on human spectra, and 0.829 on combined human-mouse spectra. The atlas is available through PTMeXchange and PRIDE, with curated site evidence integrated into UniProt and PeptideAtlas

10
Stars2Cells: Astrometric Tracking of Neurons Across Imaging Sessions

Peden-Asarch, A. M.; Honan, L. E.; Bai, J. Z.; Asarch, E. M.; Quinn, J. A.; Coffey, K. R.; Neumaier, J. F.

2026-07-08 neuroscience 10.64898/2026.07.03.736144 medRxiv
Top 0.2%
27.1%
Show abstract

Chronic calcium imaging offers a window into how single neurons and ensemble activity change across days where identifying the same neurons from one session to the next is the prerequisite for answering questions regarding learning, drift, and plasticity over time. Yet only ~2-3% of imaging laboratories publish longitudinal cross-session work, because existing registration tools depend on spatial-footprint or temporal correlations that degrade under repeated recording sessions. Here, we introduce Stars2Cells (S2C), a tracking pipeline inspired by astrometric plate-solving that represents each neuron's local geometry as a four-dimensional quad descriptor invariant to rotation, translation, and uniform scaling. S2C operates purely on centroid coordinates and combines descriptor-space matching, Random Sample Consensus (RANSAC) verification, and Hungarian assignment. Across a synthetic benchmark of 1,265 paired runs spanning 100-1,000 neurons and 8 perturbation conditions plus 1 identity sanity-floor, S2C reached pooled F1 = 98.4% compared to the standard ROI-based matching of 36.0%. To show what this enables, we applied the pipeline to dorsomedial striatum (DMS) imaging during oral fentanyl behavioral-economics self-administration. Here, we show that a conserved population-rewarded lever press response in DMS masks near-complete single-neuron turnover. This representational-drift signature we demonstrated is invisible to the bulk photometry, and resolving it requires the same-cell tracking S2C provides. S2C is distributed as a GUI-driven standalone application for both macOS and Windows, requiring no Python, command line, or virtual environment setup.

11
Quantitative Motion-Corrected PALM Links Endosome Structure and Dynamics in Live Cells

Xu, Y.; Adhikari, S.; Puchner, E. M.

2026-07-01 biophysics 10.64898/2026.06.28.735082 medRxiv
Top 0.2%
26.6%
Show abstract

Quantitative structural analysis by Photoactivated Localization Microscopy (PALM) on the nanoscale is often restricted to fixed cells because motion during prolonged data acquisition distorts image reconstruction. Here, we develop motion-corrected PALM (mcPALM), a live-cell super-resolution approach combining a conventional fluorescence channel with PALM to correct motion-induced spreading of localizations. We further introduce a photoactivation-based correction to estimate molecule numbers from incomplete trajectories. Using PI3P-marked endosomes in yeast as a dynamic model system, we show that mcPALM recovers a live-cell maturation trajectory linking motion-corrected endosome size and calibrated PI3P content, consistent with fixed-cell benchmarks. Unlike fixed-cell PALM, mcPALM preserves endosome dynamics, revealing stage-dependent directed transport and maturation-associated motility shift. Thus, mcPALM extends PALM from static structural measurements in fixed samples to integrated quantification of nanoscale structure, molecular composition and dynamics in living cells. This framework is broadly applicable to other mobile organelles and biomolecular assemblies, enabling live-cell studies on how molecular organization and dynamics are coupled to biological function.

12
V3Cell: A Vision-Guided Virtual 3D Cell Framework for Phenotypic Modeling and Perturbation Prediction

Lu, Y.; Xun, D.; chenke, X.; Xiaobo, Z.; Zhigang, Z.; Pengyu, C.; Xiwen, Y.; Zhengzheng, Y.; Jiahua, R.; Huili, H.; Jianying, H.; Pengwei, H.

2026-06-24 bioinformatics 10.64898/2026.06.23.734130 medRxiv
Top 0.2%
22.8%
Show abstract

Predicting how organoids respond to chemical perturbations is central to disease modeling and drug discovery. Existing virtual cell models operate at the single-cell level, producing static endpoint predictions from destructive assays. This leaves a critical gap at the organoid scale, where biological identity is defined by tissue-level architecture and continuous developmental dynamics rather than single-cell features. Here we introduce V3Cell, a vision-guided framework that constructs in silico surrogates of organoids directly from non-invasive bright-field microscopy. A foreground-aware model constructs static virtual 3D cells across colon, stomach, and lung organoid lineages. These virtual 3D cells closely match real samples across distributional metrics, micro-texture, and lineage-specific morphometrics, with small effect sizes for most descriptors. A temporal module further predicts developmental fate from as few as six early-frame observations and models fate-conditioned spatiotemporal trajectories that closely recapitulate real perturbation responses. V3Cell requires no omics profiling or fluorescent labeling, establishing a non-invasive brightfield-based paradigm for organoid-scale perturbation prediction. Our code and data are publicly available at https://github.com/Laineyoulu/V3Cell.

13
CryoROLE: describing large inter-domain rotation in single particle cryo-EM

Li, C.; Choi, W.; Wu, H.; Cheng, Y.

2026-07-04 biophysics 10.64898/2026.07.04.736454 medRxiv
Top 0.3%
22.1%
Show abstract

In single particle cryo-EM, analysis of continuous conformational heterogeneity has always been challenging. Both linear and deep learning-based methods treat conformational heterogeneity as perturbations to the consensus average conformation, limiting their capability in analyzing large protein motions. While classic conformational classifications are capable of handling large domain motion, they bin continuous protein dynamics into discrete static substates. Here, we present cryoROLE, a computational tool that extracts the continuous conformational dynamics embedded in the static composite map constructed from multi-body refinement into a landscape of relative orientation between the moving domains. Depicted in real space, the landscape allows intuitive interpretations of domain motion and the population of poses in the conformational space. Applying it to various biological systems reveals hidden conformational dynamics that are relevant to protein functions.

14
HyRes: Accurate Physics-Based Simulation of Dynamic Protein Structures and Interactions in Complex Environments at Scale

Li, S.; Barethiya, S.; Chen, J.

2026-06-24 biophysics 10.64898/2026.06.23.734133 medRxiv
Top 0.3%
21.7%
Show abstract

Intrinsically disordered proteins and regions (IDPs) are ubiquitous cellular regulators. Uncovering how their transient, multivalent interactions organize and fine-tune cellular processes requires transferable methods capable of deriving dynamic conformational ensembles across diverse environments at scale. Here, we present HyRes, a physics-based, hybrid-resolution protein model with atomistic backbones and intermediate-resolution sidechains that bridges the gap between atomistic accuracy and computational efficiency. Evaluated across [~]100 IDPs, HyRes generates atomistic ensembles that match or outperform state-of-the-art all-atom force fields in reproducing experimental chain dimensions, transient tertiary contacts, and local secondary structures. Demonstrating exceptional transferability, HyRes accurately captures dynamic IDP interactions in dilute phases, condensed phases, and amyloid fibril fuzzy coats. Finally, we leverage HyRes scalability to generate disordered ensembles for [~]30,000 IDPs from the human proteome and DisProt, revealing strong correlation between residual structures and cellular function and localization. HyRes and this open-access database provide unprecedented resources for IDP biology and deep learning.

15
A uniform tissue-clearing framework and mesoSPIM-ultra enable cm-scale single-neuron tracing

Pende, M.; Cregg, J. M.; Saghafi, S.; Broadbent, S.; Avdibasic, A.; Roeles, J.; Papadopoulos, S.-C.; Seaman, R. P.; Pende, N.; Mateos, M. S.; Jamwal, K.; Wunch, M.; Pasierbek, P.; Moreno-Cencerrado, A.; Korchynska, S.; Hauer, R.; Anderson, P.; Supper, P.; Kastriti, M. E.; Reumann, D.; Moorhead, M.; Graber, J. H. H.; Scholze, P.; Henschke, J. U.; Budinger, E.; Knoblich, J. A.; Klausberger, T.; Adameyko, I.; Harkany, T.; Kumar, V.; Joy, M. T.; Kiehn, O.; Dodt, H.-U.; Voigt, F.; Murawala, P.

2026-07-03 neuroscience 10.64898/2026.06.29.734841 medRxiv
Top 0.3%
21.2%
Show abstract

Tissue-clearing and light-sheet microscopy have transformed volumetric imaging of intact organs, yet limited mechanistic understanding of dehydration-based clearing continues to constrain rational protocol design and broader applicability. Here, we define the cardinal chemical and physical principles underlying dehydration-based tissue-clearing and establish a new pipeline for large-volume imaging. To maximize imaging performance, we developed the mesoSPIM-ultra, an upgraded mesoSPIM platform with a temperature-controlled sample chamber, a large field-of-view (FoV) camera and specialized optics to achieve long-working-distance, high-resolution imaging of cleared samples. We applied this approach to investigate the projectome of Chx10+ neurons, a cell population with complex axonal morphologies along the entire mouse spinal-cord and brain, and implicated in ipsilateral orienting behaviors. By combining behavioral analysis with post-hoc single-neuron reconstructions, we revealed previously inaccessible branching architectures and long-range projections extending from the brainstem to the spinal cord. Together, our work establishes a mechanistic foundation for tissue-clearing and scalable imaging.

16
zsasa: a Zig-based engine for high-throughput solvent accessible surface area at proteome scale

Nagae, T.; Tomii, K.

2026-07-03 bioinformatics 10.64898/2026.06.29.733683 medRxiv
Top 0.3%
19.5%
Show abstract

Solvent accessible surface area (SASA) is widely used to describe protein stability, ligand binding, mutation effects, and protein-protein interfaces. As structural biology workloads expand to predicted-structure collections, trajectories, and large assemblies, SASA tools must combine reproducible calculation with high throughput, low memory use, and workflow-friendly input handling. We present zsasa, a Zig-based SASA engine with command-line and Python interfaces. zsasa implements the established Shrake-Rupley and Lee-Richards algorithms, provides exact f64/f32 modes and an optional bitmask approximation, and supports batch and trajectory workflows, compressed structure inputs, and configurable atom classification including Chemical Component Dictionary (CCD)-based radii for non-standard components. In matched Shrake-Rupley validation on 4,370 Escherichia coli AlphaFold Database structures, exact double-precision zsasa reproduced FreeSASA total SASA values to near numerical identity. In 10-thread batch benchmarks on the E. coli and 23,586-structure human AlphaFold collections, zsasa was 2.94x faster than a FreeSASA batch wrapper in exact f64 mode and up to 9.70x faster in bitmask mode, with roughly 4-8x lower peak memory. Trajectory benchmarks exceeded 1,000 frames/s at tens of megabytes of peak memory, and a 4.5-million-atom PDB stress-test file completed in under five seconds. These results support zsasa as a practical tool for reproducible, low-memory generation of surface-derived structural features at large scale. zsasa is available under the MIT License at https://github.com/N283T/zsasa.

17
CellDF: Quality-controlled cell matching for whole-slide HE-IHC label transfer

Jang, E.; Huh, Y.-M.

2026-06-24 pathology 10.64898/2026.06.18.733058 medRxiv
Top 0.3%
19.1%
Show abstract

Serial-section immunohistochemistry (IHC) is the largest available source of paired hematoxylin and eosin (HE) and IHC whole slide images, yet it remains underexploited for cell-level supervision: adjacent sections sample non-identical cells, and residual registration error prevents direct assignment of IHC labels to individual HE cells. We present CellDF (Cell Displacement Field), which turns registered serial-section data into pairs of HE cells and their IHC labels by solving cell matching at whole-slide scale and assessing its reliability without ground-truth correspondences. CellDF estimates a locally adaptive residual displacement field through iterated kernel regression over each HE cells K nearest IHC candidates; a sparse-kernel variant keeps it tractable at the cell counts of a whole slide, where pairwise matchers are not. The within-tile distribution of the estimated displacements yields two ground-truth-free statistics, the directional scatter{sigma}{theta} and the between-tile angular deviation |{Delta}{theta}|, that localize matching quality more finely than landmark-based target registration error and drive a two-stage outlier filter that withholds labels where matching is unreliable. On 54 same-section HyReCo pairs,{sigma}{theta} correlates only moderately with landmark error and flags localized restaining damage that global error misses; on 30 four-marker Acrobat serial-section cases, the same statistic flags which IHC marker, if any, lies physically close enough to HE to support cell-level transfer. As a proof of concept, IHC labels transferred through CellDF trained a cell classifier on HE embeddings that generalized to held-out cells within the sample (F1 0.85, AUROC 0.88), establishing serial-section IHC as a usable cell-level labeling resource. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=78 SRC="FIGDIR/small/733058v1_ufig1.gif" ALT="Figure 1"> View larger version (42K): org.highwire.dtl.DTLVardef@a9b3dcorg.highwire.dtl.DTLVardef@15f652corg.highwire.dtl.DTLVardef@1eb3396org.highwire.dtl.DTLVardef@87dda2_HPS_FORMAT_FIGEXP M_FIG C_FIG

18
OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning

Zhang, C.; Sun, J.; Xu, Z.; Liao, R.; Yin, A.; Gao, H.; Liu, E.; Bao, Y.; Zhao, L.; Wang, G.

2026-07-12 bioinformatics 10.64898/2026.07.08.737248 medRxiv
Top 0.3%
18.9%
Show abstract

Computational modeling of cellular behavior--the virtual cell--has emerged as a stated grand challenge at the intersection of artificial intelligence and biology, yet existing foundation models remain specialized: single-cell models process dissociated transcriptomes only, spatial models require dedicated spatial-aware architectures, and perturbation predictors depend on manually curated knowledge bases that cap generalization. Here we introduce OCellus, a single nine-billion-parameter language model (Qwen3.5-9B) fine-tuned on twenty-two biological tasks that simultaneously addresses all three limitations through three coordinated technical contributions on a shared backbone. First, EvenClock encodes two-dimensional spatial coordinates as eighteen clockface sectors of text, enabling spatial reasoning on a vanilla language model without architectural modification; on ten spatial transcriptomics tasks OCellus attains 77 percent spatial-neighborhood accuracy, 96 percent spatial-cellchat accuracy, and 0.70 proportion-cosine similarity on spatial deconvolution, all without any spatial-aware architectural components. Second, per-gene language-model embeddings replace the Gene Ontology annotations that GEARS depends on, achieving Pearson correlation 0.945 on the Replogle 2022 perturbation benchmark versus 0.84 for GEARS across 457 completely unseen knockout genes. Third, OCellus-Agent provides a Planner-Router-Verifier natural-language interface that achieves 75 percent pipeline accuracy on eighty multi-task queries. Removing language-model embeddings collapses perturbation Pearson to 0.06, confirming that learned functional representations--not graph topology--drive the gain. As a cell-type encoder, OCellus ranks first among fourteen foundation models in linear-probe accuracy at 95.1 percent across four benchmark datasets, and reaches 72.6 percent average across twenty-two evaluated biological tasks--a 57-percentage-point absolute gain over the strongest baseline configuration. As a language model, OCellus uniquely generates natural-language explanations of its predictions, a capability absent from all competing methods. Code, pre-trained model weights, the graph-neural-network module, and the agent system will be made available upon publication.

19
AI-enabled reconstruction of 3D spatial multi-omics at single-cell resolution

Wang, Z.; Yan, Y.; Yang, X.; Zhang, D.; Han, C.; Zou, Q.; Du, Y.; Hu, Z.; Yuan, Z.

2026-07-14 bioinformatics 10.64898/2026.07.09.737490 medRxiv
Top 0.3%
18.8%
Show abstract

Three-dimensional (3D) spatial multi-omics provides unparalleled insights into biological activities, yet remains technically prohibitive. Here, we introduce Histo3D-MO, a hybrid experimental-computational pipeline for reconstructing single-cell-resolution 3D spatial multi-omics maps. Notably, Histo3D-MO integrates sparse, omics-disjoint spatial measurements with dense Hematoxylin and Eosin (H&E) histology through SPatial multi-Omics from h&E imaGEs (SPONGE), achieving cell-level 3D mapping across multiple omics layers. Validated using held-out slices, SPONGE substantially outperforms existing omics prediction methods. We further developed an algorithmic suite for 3D cell-type propagation and tissue-domain annotation, enabling whole-volume characterization of the tumor microenvironment. Applied to the in-house hepatocellular carcinoma data, Histo3D-MO revealed spatially organized patterns of translation efficiency, volumetric decoupling between malignant cells and monocytes, and depth-associated monocyte differentiation trajectories. Together, these results establish Histo3D-MO as a scalable framework for reconstructing single-cell-resolution 3D spatial multi-omics and interrogating tissue organization across complex biological systems.

20
Chemogenetic timestamping for the precise tracing of cell history into protein assemblies

El Hajji, L.; Gautier, A.

2026-07-10 cell biology 10.64898/2026.07.10.737712 medRxiv
Top 0.4%
18.7%
Show abstract

Self-assembling protein fibers enable to record events in single cells, bypassing the need for long-term time-lapse imaging. Fluorescent marks introduced within the growing fiber at user-defined times provide timestamps, giving access to the temporal dynamics of the recorded event. Here, we introduce CATCHFiber, a single-color timestamping strategy for tracing cellular events with high temporal resolution into self-assembling protein fibers. Relying on chemically-induced dimerization to precisely and rapidly control the incorporation of fluorescent proteins into the fiber, CATCHFiber allows the introduction of short 30-min spaced timestamps, significantly increasing the precision of event timings compared to existing methods. This increase in temporal resolution expands the use of fiber-based recorders beyond transcriptional activity, allowing to trace the kinetics of faster processes such as protein degradation, protein neosynthesis and kinase activity, and to determine the timing of cell cycle steps.